Papers with data-driven methodology

3 papers
WebNovelBench: Placing LLM Novelists on the Web Novel Distribution (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks for long-form novel generation lack scale, diversity, or objective measures.
Approach: They propose a framework that assesses long-form novel generation using an LLM-as-Judge approach.
Outcome: The proposed framework differentiates between human-written masterpieces, popular web novels, and LLM-generated content.
LSDSCC: a Large Scale Domain-Specific Conversational Corpus for Response Generation with Diversity Oriented Evaluation Metrics (N18-1)

Copied to clipboard

Challenge: Existing evaluation metrics for NRG models can't measure semantic relevance and diversity of generated results.
Approach: They propose a large-scale domain-specific conversational corpus with preprocessing and cleansing procedures for model training and a testing set for measuring the diversity of generated results.
Outcome: The proposed corpus can be taken as a new benchmark dataset for the NRG task.
Deep Supervised Contrastive Learning of Pitch Contours for Robust Pitch Accent Classification in Seoul Korean (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to classify fine-grained pitch accent patterns in Seoul Korean are limited due to variable realizations of F0 in real-world speech.
Approach: They propose a deep contrastive learning framework to classify fine-grained pitch accent patterns in Seoul Korean using a dataset of 10,093 Accentual Phrases.
Outcome: The proposed framework outperforms baseline models with state-of-the-art accuracy and F1-score.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations